Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Wang, Wenhai; Xie, Enze; Li, Xiang; Fan, Deng-Ping; Song, Kaitao; Liang, Ding; Lu, Tong; Luo, Ping; Shao, Ling

Computer Science > Computer Vision and Pattern Recognition

arXiv:2102.12122 (cs)

[Submitted on 24 Feb 2021 (v1), last revised 11 Aug 2021 (this version, v2)]

Title:Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Authors:Wenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan, Kaitao Song, Ding Liang, Tong Lu, Ping Luo, Ling Shao

View PDF

Abstract:Although using convolutional neural networks (CNNs) as backbones achieves great successes in computer vision, this work investigates a simple backbone network useful for many dense prediction tasks without convolutions. Unlike the recently-proposed Transformer model (e.g., ViT) that is specially designed for image classification, we propose Pyramid Vision Transformer~(PVT), which overcomes the difficulties of porting Transformer to various dense prediction tasks. PVT has several merits compared to prior arts. (1) Different from ViT that typically has low-resolution outputs and high computational and memory cost, PVT can be not only trained on dense partitions of the image to achieve high output resolution, which is important for dense predictions but also using a progressive shrinking pyramid to reduce computations of large feature maps. (2) PVT inherits the advantages from both CNN and Transformer, making it a unified backbone in various vision tasks without convolutions by simply replacing CNN backbones. (3) We validate PVT by conducting extensive experiments, showing that it boosts the performance of many downstream tasks, e.g., object detection, semantic, and instance segmentation. For example, with a comparable number of parameters, RetinaNet+PVT achieves 40.4 AP on the COCO dataset, surpassing RetinNet+ResNet50 (36.3 AP) by 4.1 absolute AP. We hope PVT could serve as an alternative and useful backbone for pixel-level predictions and facilitate future researches. Code is available at this https URL.

Comments:	Accepted to ICCV 2021
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2102.12122 [cs.CV]
	(or arXiv:2102.12122v2 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2102.12122

Submission history

From: Wenhai Wang [view email]
[v1] Wed, 24 Feb 2021 08:33:55 UTC (505 KB)
[v2] Wed, 11 Aug 2021 12:38:21 UTC (1,080 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Submission history

Access Paper:

References & Citations

2 blog links

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Submission history

Access Paper:

References & Citations

2 blog links

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators